Vectorize min + max and add fused minmax (#22759) - #22759
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/22759
Note: Links to docs will display an error until the docs builds have been completed. ⏳ No Failures, 124 PendingAs of commit 65a48d3 with merge base 1d95a2a ( This comment was automatically generated by Dr. CI and updates every 15 minutes. |
|
@JakeStevens has exported this pull request. If you are a Meta employee, you can view the originating Diff in D119391738. |
This PR needs a
|
Summary: Replace the iterator-returning `std::min_element` / `std::max_element` implementations of `torch::executor::vec_minf` and `vec_maxf` with four independent value-reduction lanes. A shared compile-time implementation computes only the requested extrema, with a narrowly scoped Clang vectorization hint for NEON/SSE2 targets and no fast-math requirement. Add `torch::executor::vec_minmaxf(const float* x, size_t size, float* min_out, float* max_out)` to compute both extrema together. Wire the per-tensor and both serial/parallel per-token `choose_qparams` paths to the fused helper, eliminating their separate minimum and maximum scans without changing scale/zero-point calculations. Differential Revision: D119391738
Summary: Replace the iterator-returning `std::min_element` / `std::max_element` implementations of `torch::executor::vec_minf` and `vec_maxf` with four independent value-reduction lanes. A shared compile-time implementation computes only the requested extrema, with a narrowly scoped Clang vectorization hint for NEON/SSE2 targets and no fast-math requirement. Add `torch::executor::vec_minmaxf(const float* x, size_t size, float* min_out, float* max_out)` to compute both extrema together. Wire the per-tensor and both serial/parallel per-token `choose_qparams` paths to the fused helper, eliminating their separate minimum and maximum scans without changing scale/zero-point calculations. Differential Revision: D119391738
7012c19 to
974bb46
Compare
Summary: Replace the iterator-returning `std::min_element` / `std::max_element` implementations of `torch::executor::vec_minf` and `vec_maxf` with four independent value-reduction lanes. A shared compile-time implementation computes only the requested extrema, with a narrowly scoped Clang vectorization hint for NEON/SSE2 targets and no fast-math requirement. Add `torch::executor::vec_minmaxf(const float* x, size_t size, float* min_out, float* max_out)` to compute both extrema together. Wire the per-tensor and both serial/parallel per-token `choose_qparams` paths to the fused helper, eliminating their separate minimum and maximum scans without changing scale/zero-point calculations. Differential Revision: D119391738
974bb46 to
65a48d3
Compare
Summary:
Replace the iterator-returning
std::min_element/std::max_elementimplementations oftorch::executor::vec_minfandvec_maxfwith four independent value-reduction lanes. A shared compile-time implementation computes only the requested extrema, with a narrowly scoped Clang vectorization hint for NEON/SSE2 targets and no fast-math requirement.Add
torch::executor::vec_minmaxf(const float* x, size_t size, float* min_out, float* max_out)to compute both extrema together. Wire the per-tensor and both serial/parallel per-tokenchoose_qparamspaths to the fused helper, eliminating their separate minimum and maximum scans without changing scale/zero-point calculations.Differential Revision: D119391738